Does Topical Authority Show Up in AI Citations? What 9,471 Questions Reveal

Khadija ZamanAI Search ManagerI'm Khadija Zaman, AI Search Manager at Wellows, where I lead generative and answer engine optimization (GEO/AEO) — building the automated workflows that track brand citations…Read Full Bio- We split each topic’s questions into two groups: one to measure citation coverage, and a separate group to check citations. The second group is called held-out because its questions were not used to calculate coverage.
- Coverage correlated more closely with held-out citations than the stored authority score on four of five engines: AI Overviews, AI Mode, Perplexity and ChatGPT. On Gemini the two were nearly equal, with authority slightly higher: 0.20 for coverage and 0.21 for authority.
- Citation reach on other topics had the strongest correlation of the three measures on every engine, from 0.29 to 0.37.
- Among sites grouped by outside-topic citation reach, the coverage correlation ranged from 0.208 on AI Overviews down to 0.043 on ChatGPT.
- These are associations inside one January to May 2026 window. They show where citations cluster, not what happens when you publish more pages.
Websites cited across more of a topic’s measurement questions also tended to receive more citations on separate questions in that topic. These separate questions are called held-out questions: they were kept out of the coverage calculation and used for a separate check. We saw that on all five engines Wellows tracks, across 9,471 questions and 151 topics collected between January and May 2026.
This gives us a way to examine topical authority: do sites cited across one set of questions also get cited across a separate set in the same topic? They do tend to, but the relationship varies by engine. ChatGPT had the weakest relationship between coverage and held-out citations.
The highest correlations came from how often a site was cited on other topics, which changes how the coverage numbers should be read. All of this is January to May 2026 data, collected before the July shifts we covered in our domain-authority study.
Citation coverage isn’t the same as content coverage
Citation coverage is how often a site was cited across a topic’s measurement questions. It’s something the engines do, not something you publish. A site can have deep content on a topic and almost no citation coverage, or pick up citations on questions it never wrote a dedicated page for.
That distinction decides what this study can answer. It tells you whether citation patterns carry over to other, held-out questions in the same topic, and it leaves the question of what caused them (more pages, tighter clusters, better internal links) for a different kind of test.
We also measured outside-topic reach: how often a site was cited on the other topics in the dataset. It measures how widely a site is cited across our tracked topics. It does not measure backlinks or brand awareness.
How we ran the study
The data comes from Wellows’ citation dataset for January to June 2026, covering ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode. The dataset holds 382,176 query snapshots; 86% were collected in the United States and the rest across 23 other markets. A snapshot records the results collected for a question at a particular time and in a particular market. The analysis uses the 247,397 snapshots with recorded question text, covering 9,471 English-language questions collected between January and May 2026, with 98.6% of those snapshots dated by the end of April. Questions were weighted equally despite repeated collection. We kept questions where all five engines returned at least one cited website.
Topics are project-assigned labels, pooled where projects used the same label. We kept 151 topics with at least ten distinct questions in each half.
Held-out means kept out of the coverage calculation and used for a separate check. Using the same citation observations to calculate coverage and evaluate it could make the relationship partly depend on reusing the same evidence. We therefore divided each topic’s questions into two groups before calculating coverage:
- Group A — measurement questions: used to calculate how widely each website was cited across the topic.
- Group B — held-out questions: excluded from that calculation and used to check whether the same websites were cited on different questions about the topic.
A fixed computer rule based on the topic label and question text assigned each question to a group. Within a topic, repeated observations of the same question stayed in the same group. The questions were separated, not the websites: the same website could receive citations in both groups.
Both groups came from the same January to May 2026 period. Group B was not collected later, and “held-out” does not mean an AI had never encountered those questions. It describes how we separated the observations for this analysis.

Each topic’s questions were split in two, so coverage and citations were measured on different questions.
- Coverage: for each engine, the share of measurement-half snapshots in which at least one of the other four engines cites the website, averaged across the topic’s questions. No engine is scored on its own citations.
- Outcome: each website’s citation rate on the held-out questions for the engine being tested, averaged across questions so each question has equal weight.
- Reach: the share of questions outside the topic being tested on which a website was cited at least once, using both question halves and all five engines. Sites fall into three reach groups: niche below 0.05%, mid-reach from 0.05% to below 0.5%, and widely cited at 0.5% or more.
- Authority: the 0–100 domain authority score Wellows stores with each citation, summarized using each website’s approximate median across its January to June 2026 observations. 163,156 website–topic pairs, about 82% of the sample, carry a score.
- Statistics: Spearman rank correlations calculated within each topic, with tied values given their average rank, then averaged across the 151 topics. One shuffle of held-out outcomes within topic and reach groups provides an illustrative comparison, not a statistical significance test. The similar-reach comparison converts each site’s scores to percentile ranks within its topic and reach group, then calculates one correlation across all 151 topics combined.
A payroll example: measurement versus separate check
Imagine an example website called Payroll Guide. Group A contains questions such as “What is payroll?”, “How often should employees be paid?” and “How do I fix a payroll error?” Group B contains different questions: “How do I choose payroll software?”, “What is payroll automation?” and “How do payroll tools compare?”
- Measure coverage on Group A. Suppose we are checking ChatGPT. We first look at citations from Perplexity, Gemini, AI Overviews and AI Mode. With one snapshot per question in this simplified example, if at least one of those four engines cites Payroll Guide on two of the three Group A questions, its coverage is 2 out of 3.
- Check ChatGPT on Group B. If ChatGPT cites Payroll Guide on one of the three separate questions, its held-out citation rate is 1 out of 3. None of these Group B questions contributed to the coverage calculation.
- Compare many websites within the topic. One website’s two numbers cannot establish a relationship. We check whether websites with higher Group A coverage also tend to have higher Group B citation rates, then average the topic-level correlations.
The website, questions and counts in this example are illustrative. The actual study required at least ten distinct questions in each group per topic. Where a question had repeated observations, we calculated its citation rate and averaged across questions so each question had equal weight.
Group B was held out from the coverage calculation. The outside-topic reach measure uses both question groups on other topics, as explained above. This separation tests whether citation patterns carry across questions; it does not test whether publishing more content causes more citations.
The citation test at a glance

An illustrative payroll example: measure coverage on Group A, check citations on separate Group B questions, then compare websites within the topic.
We use citations from the other four engines to measure coverage, so an engine’s own citations are not used to explain its held-out results. Engines can still favor the same sources.
The unit is a website–topic pair: 198,360 of them across 92,112 websites, each cited at least once in the measurement half. A website counts separately for each topic: a site appearing in payroll and accounting is two website–topic pairs. That means the findings apply to sites already cited in the measurement questions; they do not show how an uncited site earns its first citation.
Coverage correlated more closely than authority on four of five engines
We compared three measures against held-out citations: citation coverage, the stored authority score, and outside-topic reach. Here are the first two.

Citation coverage tracked held-out citations more closely than the authority score on four of five AI engines, with Gemini nearly equal and authority slightly higher.
| Engine | Citation coverage | Stored authority score |
|---|---|---|
| Google AI Overviews | 0.29 | 0.22 |
| Google AI Mode | 0.28 | 0.22 |
| Perplexity | 0.25 | 0.15 |
| Gemini | 0.20 | 0.21 |
| ChatGPT | 0.13 | 0.10 |
What this means: On four engines, a site’s citation coverage across a topic was more closely related to citations on other questions than its authority score was; this does not prove that publishing more pages helps.
A correlation compares how sites rank on two measures: 0.29 does not mean a 29% chance of being cited.
Coverage correlated more closely on AI Overviews, AI Mode, Perplexity and ChatGPT. The widest gap was on Perplexity, at 0.25 against 0.15. ChatGPT had the lowest values for both measures, and on Gemini the two were nearly equal, with authority slightly higher.
A higher correlation means two measured quantities moved together more closely in this sample. It does not establish a ranking factor or a causal effect. Coverage and reach use 198,360 website–topic pairs; authority uses the 163,156 pairs with a stored score. These comparisons describe the reported samples, rather than testing which measure performs better on the same sites.
Outside-topic reach had the strongest correlation
Sites cited more often across other topics also tended to receive citations on the held-out questions. Outside-topic reach had the highest correlation of the three measures on every engine.
| Engine | Outside-topic reach | Citation coverage | Stored authority score |
|---|---|---|---|
| Google AI Overviews | 0.33 | 0.29 | 0.22 |
| Google AI Mode | 0.35 | 0.28 | 0.22 |
| Perplexity | 0.29 | 0.25 | 0.15 |
| Gemini | 0.37 | 0.20 | 0.21 |
| ChatGPT | 0.29 | 0.13 | 0.10 |
What this means: Sites cited across other topics also tended to receive more citations within the topic being tested, but this does not show that publishing on unrelated subjects would help.
The gap was widest on ChatGPT and Gemini, where reach ran well ahead of coverage: 0.29 against 0.13 on ChatGPT, and 0.37 against 0.20 on Gemini.
A site may be cited within a topic partly because engines often select it across other topics too. Citation coverage alone cannot tell us how much of that visibility comes from topic expertise.
It doesn’t follow that broad publishing, PR or link building would reproduce the effect. We didn’t test any of those, and three separate correlations don’t tell you which combination would predict best in a model. Reach also uses both question halves and all five engines, so its correlation is not a like-for-like comparison with coverage measured on the first half using the other four engines.
ChatGPT had the weakest coverage signal among sites grouped by citation reach
We grouped sites within each topic into the three outside-topic citation reach bands and recalculated the coverage correlation. This is a rough comparison of citation reach, not a complete adjustment for it. We then shuffled the held-out outcomes once within those groups for an illustrative comparison; one shuffle does not establish statistical significance.

The observed correlation exceeded the shuffled value on all five engines. ChatGPT had the smallest absolute gap.
| Engine | Observed correlation | Shuffled baseline |
|---|---|---|
| Google AI Overviews | 0.208 | 0.009 |
| Google AI Mode | 0.197 | 0.008 |
| Perplexity | 0.185 | 0.008 |
| Gemini | 0.119 | 0.010 |
| ChatGPT | 0.043 | 0.012 |
What this means: Among sites grouped by citation reach, topic coverage had the weakest relationship with held-out citations on ChatGPT; it was still higher than the single shuffled comparison.
The shuffled values ranged from 0.008 to 0.012. Every observed value was higher, with the smallest absolute gap on ChatGPT: 0.043 compared with 0.012.
The data shows where ChatGPT differs, not why. Possible explanations include a preference for familiar sources, the questions tested, and how ChatGPT finds and selects sources. We would need further tests to distinguish these explanations.
What we’d do with this in a visibility program
If you’re tracking AI citations, measure published content and citation visibility separately before deciding what content to change.
- Track citations by topic and engine. One combined score can hide differences between ChatGPT, Perplexity and Google’s AI surfaces.
- Keep citation coverage separate from published coverage. Log which questions you have useful content for, independent of whether an engine cites it. Compare the two lists to find questions your content answers but where your site is not cited.
- Monitor a stable question set. Same prompts, locations and collection settings, so a change in how you measure doesn’t show up as a gain or a loss.
- Test content changes over time. Log each change, then compare later citations against a set of pages you didn’t touch.
None of this is a reason to stop building topic coverage for a ChatGPT audience. It says ChatGPT’s citations tracked coverage less closely in this sample, which is a different claim from saying depth doesn’t work there.
Wellows reports citations across ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, and keeping the engine-level view is how you tell a broad shift from one that’s confined to a single platform.
Where this fits with earlier research
Two earlier studies examined related questions. Floyi’s Topical Authority Report 2026 looked at ranked coverage in Google search results, and Kevin Indig’s analysis of Semrush data looked at category ownership in ChatGPT. Their measures differ from citation coverage, so the results are not directly interchangeable.
Wellows’ study adds a different lens: citation coverage measured directly in AI answers, tested on held-out questions, across five engines including Perplexity.
For the named-source side of the same question, our companion study of LLM topic associations looks at which sources models name by topic. It uses different data and a different outcome, so read the two side by side rather than as one result.
The question worth chasing next is what actually moves citation coverage. To find out, we need to make a specific change, measure citations afterward, and compare the results with pages we did not change.
Scope
This is an observational comparison within January to May 2026. The held-out split keeps coverage and outcomes on separate questions, and both halves come from the same window.
The sample covers questions where all five engines cited at least one website, and websites already cited in the measurement half. Pooled project labels can combine questions with different intents. The authority comparison uses the website–topic pairs that carry a stored score. No confidence intervals or repeated permutation tests are reported, so the numerical differences between engines and measures should not be treated as statistically established differences.
Our domain-authority study picked up shifts in cited-source profiles in July 2026, so a later window is the next test of whether these relationships hold.